Papers by Per Erik Solberg

3 papers
The Norwegian Dialect Corpus Treebank (2022.lrec-1)

Copied to clipboard

Challenge: The NDC Treebank consists of recordings made between 2006 and 2012 and is annotated with morphological and syntactic information.
Approach: They present the NDC Treebank of spoken Norwegian dialects in the Bokml variety of Norwegian.
Outcome: The treebank consists of 4587 speech segments, overall 66009 tokens, from 17 different Norwegian dialects from south, west, east and north of Norway.
The LIA Treebank of Spoken Norwegian Dialects (L18-1)

Copied to clipboard

Challenge: a long-term goal of this work is to develop a parser for spoken Norwegian with the immediate goal of parsing the whole LIA material.
Approach: They describe the LIA treebank of transcribed spoken Norwegian dialects and their transcription, transliteration and further morphosyntactic annotation.
Outcome: The treebank consists of 13,608 tokens, distributed over 1396 segments taken from three different dialects of spoken Norwegian.
The Norwegian Parliamentary Speech Corpus (2022.lrec-1)

Copied to clipboard

Challenge: the dataset contains recordings of meetings at the Norwegian parliament . it is the first publicly available dataset containing unscripted, Norwegian speech .
Approach: the Norwegian Parliamentary Speech Corpus is a publicly available speech dataset . it contains recordings of meetings from the Norwegian parliament with orthographic transcriptions . the dataset is intended to fill a gap in the available unscripted speech data .
Outcome: the dataset contains recordings of meetings at the Norwegian parliament with orthographic transcriptions in Norwegian Bokml and Norwegian Nynorsk.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations